Papers with high-resource languages
M-RewardBench: Evaluating Reward Models in Multilingual Settings (2025.acl-long)
Copied to clipboard
Srishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Triandi Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, Marzieh Fadaee
| Challenge: | Reward models (RMs) are primarily trained and evaluated in English and their capabilities in multilingual settings remain understudied. |
| Approach: | They construct a multilingual RM evaluation benchmark that tests the chat, safety, reasoning, and translation capabilities of RMs in 23 languages. |
| Outcome: | The proposed model performs better for high-resource languages and improves with translation quality. |
GameQA: Gamified Mobile App Platform for Building Multiple-Domain Question-Answering Datasets (2023.eacl-demo)
Copied to clipboard
Njall Skarphedinsson, Breki Gudmundsson, Steinar Smari, Marta Kristin Larusdottir, Hafsteinn Einarsson, Abuzar Khan, Eric Nyberg, Hrafn Loftsson
| Challenge: | a common problem with question-answering datasets is that they require annotators to source answers from the internet . a crowd-sourcing platform is available for low-resource languages, but it is limited in terms of information available. |
| Approach: | They propose a crowd-sourcing platform to gather multiple-domain QA data for low-resource languages. |
| Outcome: | The proposed platform rivals large QA datasets for high-resource languages in size and answerability. |
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs (2026.eacl-long)
Copied to clipboard
Manoj Balaji Jagadeeshan, Samarth Bhatia, Pretam Ray, Harshul Raj Surana, Akhil Rajeev P, Priya Mishra, Annarao Kulkarni, Ganesh Ramakrishnan, Prathosh Ap, Pawan Goyal
| Challenge: | Large language models are capable of creative generation tasks but prominently for high-resource languages. |
| Approach: | They propose to use large language models for structured poetry generation in Sanskrit . their constrained decoding method achieves 99.86% syntactic accuracy . |
| Outcome: | The proposed model outperforms the existing model in generating metrically valid Sanskrit poetry. |
MUSTS: MUltilingual Semantic Textual Similarity Benchmark (2025.acl-short)
Copied to clipboard
| Challenge: | Existing benchmarks for semantic textual similarity (STS) are limited to high-resource languages and do not include datasets annotated focusing on relatedness instead of similarity. |
| Approach: | They propose to evaluate multilingual semantic textual similarity benchmarks which span 13 languages and annotated datasets to evaluate and compare them. |
| Outcome: | The proposed method is the most comprehensive benchmark of multilingual STS methods. |
Thesis Proposal: Self-Adaptive and Epistemic Uncertainty-Guided ASR of Dense Intra-Sentential Code-Switched Speech for African Low-Resource Languages (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing multilingual and pretrained ASR systems improve general recognition accuracy but are weak at switch regions and are sensitive to language imbalance during adaptation. |
| Approach: | They propose a self-adaptive and epistemic uncertainty-guided framework for African low-resource code-switched ASR using Hausa–English and Hausa-Yorùbá as case studies. |
| Outcome: | The proposed framework is based on Hausa–English and Hausa-Yorùbá as case studies. |
Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity (2024.findings-eacl)
Copied to clipboard
Eric Khiu, Hasti Toossi, Jinyu Liu, Jiaxu Li, David Anugraha, Juan Flores, Leandro Roman, A. Seza Doğruöz, En-Shiun Lee
| Challenge: | Existing approaches for predicting the performance of NLP models for low-resource languages (LRLs) focus on high-resourced languages, overlooking LRLs and domain shifts. |
| Approach: | They investigate the impact of domain similarity on predicting performance of machine translation models in low-resource languages. |
| Outcome: | The results show that domain similarity has the most important impact on predicting the performance of Machine Translation models. |
A Testset for Context-Aware LLM Translation in Korean-to-English Discourse Level Translation (2025.coling-main)
Copied to clipboard
| Challenge: | Recent studies indicate that for high-resource languages, LLM surpasses encoder-decoder neural machine translation (NMT) models. |
| Approach: | They propose to construct a Korean-English discourse-level corpus with 600 text instances featuring six linguistic phenomena: lexical ambiguity, zero anaphora, slang, idiom, figurative language, and implicature. |
| Outcome: | The proposed corpus of 600 text instances features six linguistic phenomena, including lexical ambiguity, zero anaphora, slang, idiom, figurative language, and implicature. |
AfriVox: Probing Multilingual and Accent Robustness of Speech LLMs (2026.eacl-long)
Copied to clipboard
Busayo Awobade, Mardhiyah Sanni, Tassallah Abdullahi, Chibuzor Okocha, Kelechi Ezema, Devendra Deepak Kayande, Lukman Enegi Ismaila, Tobi Olatunji, Gloria Ashiya Katuka
| Challenge: | Recent advances in multimodal and speech-native large language models have delivered impressive speech recognition, translation, understanding, and question-answering capabilities for high-resource languages. |
| Approach: | They propose to benchmark African languages and African-accented French, Arabic, and 100+ African English accents across 20 African languages. |
| Outcome: | The proposed model outperforms traditional speech transcription and translation models in African languages and non-native French or English accents. |
Breaking Down Multilingual Machine Translation (2022.findings-acl)
Copied to clipboard
| Challenge: | Multilingual training is an essential ingredient in machine translation systems . but it has different effects in different multilingual settings, such as many-to-one, one-tomany and many- to-many learning . |
| Approach: | They compare multilingual training settings with encoders and decoders initialized by multilingual learning . they find important attention heads for each language pair and compare their correlations during inference . |
| Outcome: | The proposed models outperform the best models for high-resource languages and one-to-many models for low-resourced languages. |
Why Can’t Discourse Parsing Generalize? A Thorough Investigation of the Impact of Data Diversity (2023.eacl-main)
Copied to clipboard
| Challenge: | Discourse parsing performance is not reliable for high-resource languages such as English . a heterogeneous training regime is critical for stable and generalizable models . |
| Approach: | They investigate the impact of genre diversity on RST parsing stability . they use two largest RST corpora of English with text from multiple genres . |
| Outcome: | The proposed model can generalize to text types unseen during training, but it is not reliable for high-resource languages. |
Towards Making the Most of ChatGPT for Machine Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Prior studies have shown that ChatGPT achieves comparable results to commercial systems for high-resource languages, but lags behind in complex tasks, e.g., low-resourced and distant-language-pairs translation. |
| Approach: | They propose task-specific prompts and domain-specific prompts which are based on task information and domain information and a task-specific prompt. |
| Outcome: | The proposed prompts improve the performance of ChatGPT in complex tasks and generate hallucinations for non-English-centric tasks. |
CHIA: CHoosing Instances to Annotate for Machine Translation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Neural machine translation systems perform poorly on low-resource language pairs, for which large-scale parallel data is unavailable. |
| Approach: | They propose a method for selecting instances to annotate for machine translation using existing multi-way parallel datasets. |
| Outcome: | The proposed method outperforms unsupervised methods on 20 languages and a multi-way parallel dataset on high-resource languages. |
LaoPLM: Pre-trained Language Models for Lao (2022.lrec-1)
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) can capture different levels of concepts in context . previous work on Lao has been hampered by the lack of annotated datasets . |
| Approach: | They construct a text classification dataset to alleviate the resource-scarce situation of Lao . they evaluate them on two downstream tasks: part-of-speech tagging and text classification . |
| Outcome: | The proposed model can capture different levels of concepts in context and generate universal language representations. |
Error Analysis of Uyghur Name Tagging: Language-specific Techniques and Remaining Challenges (L18-1)
Copied to clipboard
| Challenge: | despite efforts at name tagging, there is limited understanding on the performance ceiling . despite the high-resource language, there are very few natural language processing tools available . |
| Approach: | They propose to use a machine learning model to identify Uyghur name tagger errors . they conclude that such a model is unlikely to be effective for Uygur, or low-resource languages . |
| Outcome: | The proposed model is unlikely to be effective for Uyghur, or low-resource languages in general, the authors argue . they show that the proposed model can be used for high-res languages with superficial features . |
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual large language models (LLMs) exhibit factual inconsistencies across languages . authors identify two primary sources of error: insufficient engagement of reliable English-centric mechanism for factual recall, and incorrect translation from English back into the target language for the final answer. |
| Approach: | They propose two vector interventions to redirect the model toward better internal paths for higher factual consistency. |
| Outcome: | The proposed interventions increase the recall accuracy by over 35 percent for the lowest-performing language. |
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment. |
| Approach: | They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions. |
| Outcome: | The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages. |
Using Convolution Neural Network with BERT for Stance Detection in Vietnamese (2022.lrec-1)
Copied to clipboard
| Challenge: | Stance detection is a task of automatically eliciting stance information towards a specific claim made by a primary author. |
| Approach: | They propose an architecture using transformers to detect stances in Vietnamese claims . they exploit BERT to extract contextual word embeddings instead of traditional word2vec models . |
| Outcome: | The proposed model outperforms the previous methods on a public dataset. |
Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent speech foundation models excel at multilingual automatic speech recognition (ASR) for high-resource languages, but their performance drops substantially on low-resourced languages due to the limited data availability. |
| Approach: | They propose a Depth-Aware Model Adaptation framework that allocates adaptation capacity according to each layer’s role. |
| Outcome: | The proposed framework matches or surpasses state-of-the-art accuracy with 80% fewer trainable parameters and achieves 29% error reduction under extreme data scarcity. |
Large Language Models for Multilingual Previously Fact-Checked Claim Detection (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new study evaluates large language models for multilingual previously fact-checked claim detection . authors assess seven LLMs across 20 languages in monolingual and cross-lingual settings . |
| Approach: | They evaluate large language models for multilingual previously fact-checked claim detection . they find they perform well for high-resource languages, struggle with low-resourced languages . |
| Outcome: | The proposed model performs well for high-resource languages, but struggle with low-resourced languages. |
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing LLMs consistently underperform across all tasks, with 10-shot learning and fine-tuning offering only limited improvements. |
| Approach: | They introduce FormosanBench, a benchmark for evaluating LLMs on low-resource Austronesian languages. |
| Outcome: | The proposed benchmark covers three endangered Formosan languages: Atayal, Amis, and Paiwan . existing LLMs consistently underperform across all tasks, with 10-shot learning and fine-tuning offering only limited improvements. |
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit impressive zero/few-shot inference and generation quality for high-resource languages (HRLs). |
| Approach: | They propose to reduce the cost of processing LRLs by code-mixing, translation, and transliteration of LRL to HRLs to ensure that predictive and generative qualities are not compromised. |
| Outcome: | The proposed model reduces the cost of processing LRLs while ensuring that predictive and generative qualities are not compromised. |
Automatic Transcription of Handwritten Old Occitan Language (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to handwritten text recognition have shown promising results, but low-resource languages often lack resources. |
| Approach: | They propose an HTR approach that leverages the Transformer architecture for recognizing handwritten Old Occitan language. |
| Outcome: | The proposed approach surpasses state-of-the-art models for Old Occitan HTR, including open-source Transformer-based models and commercial applications like Google Cloud Vision. |
Enhancing Multilingual RAG Systems with Debiased Language Preference-Guided Query Fusion (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies show that mRAGs exhibit a perceived preference for high-resource languages, particularly English. |
| Approach: | They propose a debiased language preference metric to explicitly factor out structural priors . they propose mRAG framework that leverages monolingual alignment to optimize cross-lingual retrieval and generation. |
| Outcome: | The proposed framework outperforms baselines for English pivoting and mRAG in multiple languages. |
TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMs (2026.acl-long)
Copied to clipboard
Ricardo Rei, Nuno M Guerreiro, José Pombal, João Alves, Amin Farajian, Pedro Teixeirinha, Andre Martins
| Challenge: | Large Language Models (LLMs) are emerging as the de facto solution for multilingual machine translation. |
| Approach: | They propose a suite of LLMs that can be fine-tuned to deliver strong performance on translation and multilingual general-purpose text capabilities. |
| Outcome: | The proposed models outperform existing models on translation and general-purpose tasks. |
Towards Language-Agnostic STIPA: Universal Phonetic Transcription to Support Language Documentation at Scale (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing ASR systems focus on orthographic output for high-resource languages, but STIPA can be used as a language-agnostic interface for documenting under-resourced and unwritten languages. |
| Approach: | They propose to use the International Phonetic Alphabet (STIPA) to generate phonetic transcriptions using a language-agnostic interface. |
| Outcome: | The proposed model reduces phonetic error rates even in low-resource settings and can be used for documenting under-resourced and unwritten languages. |
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages. (2026.findings-acl)
Copied to clipboard
Tassallah Abdullahi, Macton Mgonzo, Mardiyyah Oduwole, Paul Okewunmi, Abraham Toluwase Owodunni, Ritambhara Singh, Carsten Eickhoff
| Challenge: | Current guardian models are predominantly Western-centric and optimized for high-resource languages . low-resourced African languages are vulnerable to evolving harms, cross-lingual failures, cultural misalignment . |
| Approach: | They propose a policy-based safety benchmark for African languages built from adversarial queries authored by 155 domain experts across sensitive fields. |
| Outcome: | The proposed model overestimates multilingual safety, cross-lingual transfer provides partial but insufficient coverage, and dynamic models struggle to localize African-language contexts. |
One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | a recent study has focused on setups that favor romanization for cross-lingual transfer . a fidelity-based approach is needed to improve performance for high-resource languages . |
| Approach: | They propose to pretrain LMs from scratch on romanized and original texts for six languages . they find that romanization improves encoding efficiency for segmental scripts at a negligible cost . |
| Outcome: | The proposed method reduces the loss of script-specific information and dilution of language-specific representations from increased subword overlap. |
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing LLMs either reason in English and translate, or simply fail on multi-step Bengali math. |
| Approach: | They propose a Bengali mathematical reasoning model called GanitLLM with a difficulty-aware Bengali math corpus and a curriculum-based GRPO pipeline. |
| Outcome: | The proposed model improves on Bn-MGSM and Bn MSVAMP by +8 and +7 accuracy points while increasing the percentage of Bengali reasoning tokens from 14% to over 88% and reducing solution length from 943 to 193 words. |